npj Antimicrobials and Resistance
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match npj Antimicrobials and Resistance's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Hessel, M.; Inda Diaz, J. S.; Sjöberg, A.; Salva-Serra, F.; Helldal, L.; Jirstrand, M.; Johnning, A.; Kristiansson, E.; Skovbjerg, S.
Show abstract
Antimicrobial resistance is a public health challenge, driving the need for rapid, cost-effective diagnostic support tools. Artificial intelligence (AI) may enable prediction of susceptibility to untested antibiotics from known susceptibility results, but prospective clinical validation is required before routine use. We evaluated an AI-based decision support method, trained on invasive isolates from the European Surveillance System (TESSy), for prediction of antibiotic susceptibility in clinical Escherichia coli urine isolates. The evaluation included 99 E. coli isolates from urine samples with diversity in age, sex, and antibiotic susceptibility. Predictions were evaluated for 14 antibiotics using patient metadata and susceptibility results for 4-8 antibiotics as input. Prediction uncertainty was handled using conformal prediction, allowing abstention when confidence was insufficient. EUCAST disk diffusion test results were used as reference and genomic sequence data was used to explore mechanisms of the AI performance. Without conformal prediction, 84% of predictions were correct when susceptibility results of six antibiotics were used to predict susceptibility to eight additional antibiotics. Across all predictions generated using susceptibility results for six antibiotics as input, the major and very major error rates were 19% and 12%, respectively. Prediction errors varied between antibiotics and were associated with certain phenotypic and genotypic resistance patterns. Conformal prediction reduced errors but increased abstentions; at confidence levels of 90%, 95%, and 97.5%, the model abstained in 9.6%, 14%, and 22% of instances. The method showed promising performance, but its clinical use remains limited and may require diagnostic data beyond susceptibility test results and demographic variables.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Santoyo, G.; Flores, A.; Castelan-Sanchez, H. G.; Valenzuela-Ruiz, V.; de los Santos-Villalobos, S.; Mitra, D.; Babalola, O. O.; Schoebitz, M.; Orozco-Mosqueda, M. d. C.
Show abstract
Plant growth-promoting bacterial endophytes represent a sustainable strategy for enhancing agricultural productivity while reducing reliance on synthetic fertilizers and pesticides. This study focused on the genomic and functional characterization of two endophytic bacterial strains, R11F and R19M, isolated from bean and maize roots, respectively. Comparative analyses based on 16S rRNA gene sequences, average nucleotide identity (ANI), and genome-to-genome distance calculations (GGDC) classified both isolates as Pseudomonas palleroniana. Comparative genomic analyses revealed highly conserved genomes containing genes associated with plant colonization, phosphate solubilization, stress adaptation, heavy metal resistance, and hydrocarbon degradation. Genome mining further identified 17 and 18 biosynthetic gene clusters (BGCs) in R11F and R19M, respectively, including non-ribosomal peptide synthetases (NRPS), pyoverdine, NRP-metallophores, RiPP-like compounds, arylpolyenes, {beta}-lactones, terpenes, NAGGN, and hydrogen cyanide. Strain-specific BGCs associated with syringomycin and viscosin biosynthesis were identified in R11F, whereas R19M harbored clusters related to asplenin and kolossin biosynthesis. In vitro assays confirmed indole production, phosphate solubilization, and siderophore production, as well as the ability of both strains to grow in nitrogen-free medium. Both strains significantly inhibited the growth of Fusarium oxysporum, Phytophthora cinnamomi, and Colletotrichum gloeosporioides. Furthermore, plant inoculation assays demonstrated host-dependent growth promotion, with R11F showing the most consistent improvements in plant growth parameters in tomato, wheat, and lentil. Overall, the integration of comparative genomics and experimental validation demonstrates that P. palleroniana R11F and R19M possess complementary traits associated with plant growth promotion, pathogen suppression, saline stress adaptation, and bioremediation.
Omani, R.; Maina, G. N.; Fasina, F. O.
Show abstract
Public genomic repositories can support antimicrobial resistance (AMR) surveillance, but unequal sampling can bias interpretation. We characterised AMR determinants, multicountry genomic cluster overlap and surveillance gaps across Africa using an NCBI Pathogen Detection snapshot retrieved on 24 August 2026 for 55 African Union member states. Records were validated and deduplicated by BioSample, and complete AMRFinderPlus calls were summarised across five United Nations M49 subregions and eight overlapping regional economic communities (RECs). Country-pair cluster overlap was assessed using the Jaccard index, while project-based and composition-standardised sensitivity analyses evaluated repository bias. The dataset contained 86,829 unique BioSamples from 51 states; South Africa, Malawi and Kenya contributed 55.8%. Complete extended-spectrum {beta}-lactamase calls were detected in 21,513 isolates and carbapenemase calls in 4,642. blaCTX-M-15 dominated the ESBL profile, while NDM and OXA types predominated. Seventy clusters contained carbapenemase-positive isolates from at least two countries. A shared REC covered all participating countries in 38 clusters, while 32 crossed REC boundaries. Normalised country-pair overlap was low, with a maximum Jaccard index of 9.5%. Project balancing reduced the Northern African carbapenemase estimate from 32.3% to 17.9% and the Eastern African ESBL estimate from 36.9% to 12.5%. Public repositories identify determinants and clusters for investigation but do not estimate prevalence or transmission. AMR surveillance should combine national confirmation, regional institution-led investigation where countries share an REC, and continent-wide coordination through Africa CDC for cross-REC signals, supported by representative One Health sampling, standardised metadata and sustained African sequencing capacity.
Zhao, C.; Ji, Z.
Show abstract
Spatial domain detection is a central task in spatial transcriptomics, yet existing methods exhibit highly variable performance across datasets. We introduce L-STAR, a visual LLM-guided, consensus-based framework that leverages the visual reasoning capacity of large language models to adaptively rank and integrate spatial domain detection methods. L-STAR achieves robust and consistently improved performance, outperforming single spatial domain detection methods across diverse datasets.
Pham, T. M.; Smith, J. T.; Mortimer, T. D.; Grad, Y.; Earl, A. M.; Lewis, I. A.; PRIME Consortium,
Show abstract
Background Using a population-based cohort from the Calgary Health Zone (CHZ), Canada, we integrated longitudinal antimicrobial susceptibility and prescribing data with the whole genome sequences of five major pathogens. We aimed to assess how antimicrobial resistance (AMR) responds to prescribing changes and determine which bacterial strains shape these dynamics. Methods We analysed antibiotic prescribing rates, clinical and genomic data from 7,271 Staphylococcus aureus, 1,609 Enterococcus faecalis, 801 Enterococcus faecium, 11,363 Escherichia coli, and 2,319 Klebsiella pneumoniae isolates, associated with bacteraemia episodes in the CHZ between 2006-2022. Genomic clusters (referred to as strains) were identified using StrainGST and assigned to known sequence types (STs) or clonal complexes (CCs). Strain-level incidence, stratified by community-onset (isolates collected [≤]48h after admission) and hospital-onset (>48h after admission), AMR phenotypes, and prescribing rates were modelled using negative-binomial and binomial regression. Temporal trends were quantified using average annual percentage change (AAPC). Findings Between 2010-2022, fluoroquinolone prescribing declined in both community (AAPC=-6.8% [95% CI -8.1, -5.4]; p<0.0001) and hospital settings (AAPC=-5.1% [-6.5, -3.7]; p<0.0001). This was accompanied by a significant reduction in fluoroquinolone resistance among Gram-positive species. Specifically, S aureus bacteraemia resistant to clinically important antibiotics, cloxacillin, ciprofloxacin, erythromycin, and clindamycin, declined from 2006 to 2022, mostly in hospital-onset cases (AAPC=-16.0%, [-19.3%, -12.7%], p<0.0001). In E coli, ceftriaxone and ciprofloxacin resistance were clustered in ST131 and the emerging ST1193; the latter increased steadily, particularly in community-onset cases (AAPC=17.7%, [0.0%, 30.0%], p<0.0001). CTX-M-27-producing E coli ST131 strains increased (AAPC=23.8%, [17.4%, 30.5%], p<0.0001) between 20082022, while CTX-M-14-producing E coli ST131 declined (AAPC=-15.9%, [-21.3%, -10.2%], p<0.0001) between 2013-2022. These trends were paralleled by an increase in community cephalosporin prescribing (AAPC=7.3%, [4.2%, 10.5%], p<0.0001) between 2010-2022. For K pneumoniae, hypervirulent ST23 was most common (N=88) with an increasing trend in incidence (AAPC=3.0%, [-2.8%, 9.2%]) between 2006-2019. Conclusions The contrasting resistance trends between Gram-positive and Gram-negative species underscore the complexity of AMR control efforts. Effective strategies will require stewardship efforts targeting multiple drug classes, genomic surveillance for emerging resistant strains, and interventions extending beyond hospital settings.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.
Show abstract
Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.
Show abstract
Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.
Bingham, J. C.; Arussy, N.
Show abstract
Active Feature Acquisition (AFA) adaptively selects which diagnostic test to order next and offers a route to reduce unnecessary laboratory testing in acute care. Existing clinical AFA evaluations, however, assume every feature can be retrieved on demand and split data at the visit level, both of which inflate apparent performance. We re-evaluate cost-aware AFA under constraints designed to reflect deployment. From MIMIC-IV we constructed a cohort of 64,766 acute admissions (39,884 patients; 21 conditions; 55 features in 30 test panels) with a patient-level split, a 12-hour decision cutoff, and a per-patient availability mask from what was actually measured, and priced panels using the 2026 Medicare fee schedule under panel-level billing. We evaluated EIG-Cost, which scores each panel by Monte-Carlo Expected Information Gain penalised by its dollar cost, against eight published methods across budgets \30--$60 over five patient-level resamples. At a $30 budget, EIG-Cost achieved the highest macro-F1 (0.188, 95% CI [0.185, 0.191]) at the lowest cost ($17.28), exceeding the strongest baseline in all five resamples (p<0.001; Cohen's d=4.0), and led at every budget. Three of the eight methods collapsed to a vitals-only baseline (macro-F1 approx 0.040), acquiring nothing even at higher budgets, a genuine failure to adapt to availability rather than a budget limitation. Despite modest absolute accuracy, EIG-Cost's probabilities were well-calibrated (expected calibration error $0.048$). Under realistic availability constraints, clinical AFA is substantially harder than full-availability benchmarks imply, several published methods fail outright, and cost-aware information-gain scoring is a robust choice in this harder setting.
Mukherjee, E. M.; Asiaee, A.; Park, D.; Krantz, M. S.; Stone, C. A.; Martin-Pozo, M.; Phillips, E. J.
Show abstract
Importance: Immune checkpoint inhibitors (ICIs) produce diverse immune toxicities, but whether checkpoint blockade also modifies associations between other drugs and adverse events is poorly understood. Objective: To define ICI-associated toxicity organization and determine whether drug-associated adverse events and onset vary with ICI exposure and checkpoint pathway. Design and Setting: Cross-sectional analysis of deduplicated FAERS reports from 2016 through 2025; analyses performed in 2026. Participants: Among 13,701,106 deduplicated reports, 2,365,269 were cancer associated and 256,940 contained an ICI. Median age among cancer reports with observed age was 66 years (IQR, 56-75 years); 1,031,999 (43.6%) were female and 1,003,154 (42.4%) were male. Exposures: ICI exposure in any reported drug role, individual primary-suspect drugs, and checkpoint-pathway exposure. Main Outcomes and Measures: Reporting odds ratios (ORs), cross-organ adverse-event communities, adjusted primary-suspect drug x ICI interaction ORs for Stevens-Johnson syndrome/toxic epidermal necrolysis (SJS/TEN), drug reaction with eosinophilia and systemic symptoms (DRESS), acute generalized exanthematous pustulosis (AGEP), interstitial nephritis, drug-induced liver injury (DILI), and vomiting (VOM), and accelerated failure-time model time ratios for documented onset. Results: Of 3001 eligible Preferred Terms in cancer-associated reports, 2091 differed at a false discovery rate (FDR) less than .05. Four cross-organ toxicity communities were identified. Of 138 eligible drug-phenotype pairs, 65 had FDR-significant interactions, including moxifloxacin-SJS/TEN amplification (interaction OR, 101.72; 95% CI, 39.11-264.55), enfortumab vedotin-SJS/TEN attenuation (interaction OR, 0.17; 95% CI, 0.13-0.23), and omeprazole-interstitial nephritis amplification (interaction OR, 10.35; 95% CI, 7.62-14.05). Among 60,324 reports contributing to temporal analyses, ICI exposure was associated with longer adjusted documented time to onset for 5 of 6 phenotypes (time ratios, 1.37-1.59) but not AGEP (time ratio, 0.99; 95% CI, 0.67-1.46). Temporal associations also differed across checkpoint pathways. Conclusions and Relevance: ICIs were associated with a structured cross-organ toxicity landscape, phenotype-specific modification of drug-associated adverse events, and distinct temporal patterns across checkpoint pathways. These findings support checkpoint blockade as a modifier of drug-associated toxicity and motivate longitudinal and mechanistic validation.
Xuan, H.; Huang, Y.; Bian, J.
Show abstract
Machine-learning models of the human microbiome are trained mostly on stool samples from single cohorts, limiting cross-body-site representation and cross-study generalization. Progress is constrained less by algorithms than by the absence of a harmonized multi-body-site corpus carrying the technical metadata needed to model, rather than ignore, batch structure. Here we release Corpusome, a harmonized two-tier cross-body-site human microbiome corpus for representation learning: a harmonized corpus of 187,546 human microbiome samples integrating standardized profiles from curatedMetagenomicData, the American Gut Project, and the EBI MGnify platform. Corpusome follows a two-tier design preserving both functional depth and cross-body-site breadth: a shotgun tier (22,588 samples, 93 studies) with species- and pathway-level profiles, and a 16S tier (164,958 samples, from a full pull of 708 MGnify studies) with genus-level profiles extending coverage to oral, skin, respiratory, and urogenital sites. It spans six body sites and two modalities, with harmonized metadata for batch-aware modelling. Body-site signal exceeds technical/source variance in the 16S tier by approximately 2.4-fold.
Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.
Show abstract
Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
McPhillips, C. H.; Reilly, E. T.; Stolberg-Mathieu, G.; Nielsen, K.; Gottlieb, A. D.; Madjarov, G.; Roager, H. M.; Nielsen, D. S.; Krych, L.
Show abstract
Next-generation sequencing (NGS) of the prokaryotic 16S rRNA gene revolutionized gut microbiome research two decades ago. However, short read lengths remain an inherent limitation of platforms such as the widely used Illumina platforms (2 x 150-300 bp). Recent advances in Oxford Nanopore Technologies (ONT) flow cell chemistry (R10.4.1) have substantially improved sequencing accuracy. Combined with a custom multiple-primer strategy that comprehensively targets 16S rRNA gene variants to generate near-full-length amplicons, this approach enables read-by-read taxonomic classification, a feature not feasible with short-read sequencing platforms. Although our multiple-primer strategy could enable parallel sequencing of more than 18,000 samples (192 x 96), current flow cell capacity offers sufficient sequencing depth for approximately 1,000-1,500 samples. To validate the scalability and our per-read classification pipeline, we show that more than a thousand human fecal microbiome samples spiked with two bacterial strains (Imtechella halotolerans and Allobacillus halotolerans), not otherwise present in human fecal samples, can be successfully sequenced on a single flow cell, achieving a per-molecule error rate sufficient for direct per-read classification and at an adequate read depth for downstream analysis. This level of scalability significantly reduces per-sample costs, making the approach more accessible to a broader research community. To embrace these advancements, we have developed RubyRed, a pipeline that processes raw sequencing data and assigns taxonomic classifications on a per-read basis. Using spike-in references (I. halotolerans and A. halotolerans), we demonstrate high mean single-read sequencing accuracy (99% and 98.9%, respectively), with the majority of reads exceeding the canonical threshold required for species-level taxonomic classification based on the 16S rRNA gene.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.